Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos

نویسندگان

  • Mingze Xu
  • Chenyou Fan
  • Yuchen Wang
  • Michael S Ryoo
  • David J Crandall
چکیده

In a world in which cameras are becoming more and more pervasive, scenes in public spaces are often captured from multiple perspectives by diverse types of cameras, including surveillance and wearable cameras. An important problem is how to organize these heterogeneous collections of videos by finding connections between them, such as identifying common correspondences between people both appearing in the videos and wearing the cameras. In this paper, we consider scenarios in which multiple cameras of different types are observing a scene involving multiple people, and we wish to solve two specific, related problems: (1) given two or more synchronized third-person videos of a scene, produce a pixel-level segmentation of each visible person and identify corresponding people across different views (i.e., determine who in camera A corresponds with whom in camera B), and (2) given one or more synchronized third-person videos as well as a first-person video taken by a wearable camera, segment and identify the camera wearer in the third-person videos. Unlike previous work which requires ground truth bounding boxes to estimate the correspondences, we jointly perform the person segmentation and identification. We find that solving these two problems simultaneously is mutually beneficial, because better fine-grained segmentations allow us to better perform matching across views, and using information from multiple views helps us perform more accurate segmentation. We evaluate our approach on a challenging dataset of interacting people captured from multiple wearable cameras, and show that our proposed method performs significantly better than the state-of-the-art on both person segmentation and identification.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Actor and Observer: Joint Modeling of First and Third-Person Videos

Several theories in cognitive neuroscience suggest that when people interact with the world, or simulate interactions, they do so from a first-person egocentric perspective, and seamlessly transfer knowledge between third-person (observer) and first-person (actor). Despite this, learning such models for human action recognition has not been achievable due to the lack of data. This paper takes a...

متن کامل

CRF-Based Context Modeling for Person Identification in Broadcast Videos

2 We are investigating the problem of speaker and face identification in broadcast videos. 3 Identification is performed by associating automatically extracted names from overlaid texts 4 with speaker and face clusters. We aimed at exploiting the structure of news videos to solve 5 name/cluster association ambiguities and clustering errors. The proposed approach combines 6 iteratively two Condi...

متن کامل

Trajectory aligned features for first person action recognition

Egocentric videos are characterised by their ability to have the first person view. With the popularity of Google Glass and GoPro, use of egocentric videos is on the rise. Recognizing action of the wearer from egocentric videos is an important problem. Unstructured movement of the camera due to natural head motion of the wearer causes sharp changes in the visual field of the egocentric camera c...

متن کامل

Cross-View Person Identification by Matching Human Poses Estimated with Confidence on Each Body Joint

Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interests recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does n...

متن کامل

Social Behavior Prediction from First Person Videos

This paper presents a method to predict the future movements (location and gaze direction) of basketball players as a whole from their first person videos. The predicted behaviors reflect an individual physical space that affords to take the next actions while conforming to social behaviors by engaging to joint attention. Our key innovation is to use the 3D reconstruction of multiple first pers...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2018